Papers with layer representations
Multilingual, Multi-scale and Multi-layer Visualization of Intermediate Representations (D19-3)
Copied to clipboard
| Challenge: | Currently, the main alternatives to deal with sequences are Recurrent Neural Networks (RNN) architectures and the Transformer. |
| Approach: | They propose a web-based tool that visualizes the sentence and token representations of RNNs and Transformer architectures at the sentence level. |
| Outcome: | The proposed visualization tool analyses gender inequalities in contextual word embeddings and the common language representation in a multilingual machine translation system. |
Quantifying the Contextualization of Word Representations with Semantic Class Probing (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Pretrained language models are effective in solving NLP tasks, but there are still questions about how and why they work so well. |
| Approach: | They use BERT to quantify contextualization by studying the extent of inference . they show that top layer representations support highly accurate inference of semantic classes . |
| Outcome: | The proposed model is highly accurate, but weak in the lower layers . it is more task-specific after finetuning while lower layers are more transferable . |
Choose Your Transformer: Improved Transferability Estimation of Transformer Models on Classification Tasks (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing models for NLP tasks require fine-tuning, but it is computationally infeasible. |
| Approach: | They propose an approach that inexpensively estimates a ranking of the expected performance of a given set of transformer language models for a specific task. |
| Outcome: | The proposed model improves the Pearson correlation coefficient between the true model ranks and the estimate. |
Enhancing Chain-of-Thought Reasoning with Critical Representation Fine-tuning (2025.acl-long)
Copied to clipboard
Chenxi Huang, Shaotian Yan, Liang Xie, Binbin Lin, Sinan Fan, Yue Xin, Deng Cai, Chen Shen, Jieping Ye
| Challenge: | Representation Fine-tuning (ReFT) is a proposed method for improving parameter efficiency . however, it yields suboptimal performance, as fixed-position representations have uncertain impact on outputs . |
| Approach: | They propose a method that fine-tunes critical representations in a low-rank linear subspace while freezing the base model. |
| Outcome: | The proposed method improves accuracy of LLaMA-2-7B and ReFT by 18.2 and 3.8 on GSM8K. |